Àá½Ã¸¸ ±â´Ù·Á ÁÖ¼¼¿ä. ·ÎµùÁßÀÔ´Ï´Ù.
KMID : 1022420190110030039
Phonetics and Speech Sciences
2019 Volume.11 No. 3 p.39 ~ p.47
Performance of Korean spontaneous speech recognizers based on an extended phone set derived from acoustic data
Bang Jeong-Uk

Kim Sang-Hun
Kwon Oh-Wook
Abstract
We propose a method to improve the performance of spontaneous speech recognizers by extending their phone set using speech data. In the proposed method, we first extract variable-length phoneme-level segments from broadcast speech signals, and convert them to fixed-length latent vectors using an long short-term memory (LSTM) classifier. We then cluster acoustically similar latent vectors and build a new phone set by choosing the number of clusters with the lowest Davies-Bouldin index. We also update the lexicon of the speech recognizer by choosing the pronunciation sequence of each word with the highest conditional probability. In order to analyze the acoustic characteristics of the new phone set, we visualize its spectral patterns and segment duration. Through speech recognition experiments using a larger training data set than our own previous work, we confirm that the new phone set yields better performance than the conventional phoneme-based and grapheme-based units in both spontaneous speech recognition and read speech recognition.
KEYWORD
acoustic units, phone set, spontaneous speech recognition, broadcast data
FullTexts / Linksout information
Listed journal information
ÇмúÁøÈïÀç´Ü(KCI)